Back

Nature Methods

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Nature Methods's content profile, based on 385 papers previously published here. The average preprint has a 0.41% match score for this journal, so anything above that is already an above-average fit.

1
ZenReg: A modular Python platform for fast and memory-efficient N-dimensional microscopy image registration

Musacchio, F.; Fuhrmann, M.

2026-08-13 neuroscience 10.64898/2026.08.07.743572 medRxiv
Top 0.1%
52.5%
Show abstract

Motion artifacts are almost unavoidable in functional time-lapse and structural volumetric multiphoton microscopy. They arise from respiration, heartbeat, locomotion, awake behavior, instrument heating, mechanical vibration, and slow drift, while the recorded signal is often photon-limited, blurred by scattering, and biologically time varying. Consequently, motion correction is frequently an essential prerequisite for quantitative bioimage analysis rather than a merely cosmetic preprocessing operation. Edge- and landmark-centric registration strategies are often poorly matched to these data because useful structures may be sparse, diffuse, out-of-focus, or changing in fluorescence intensity. We present ZenReg, an open-source Python platform that formulates common 2D+t, 3D, and 3D+t microscopy registration tasks as a modular family of geometry-preserving alignment problems. ZenReg combines Fourier phase correlation, intensity-based StackReg-style alignment, NoRMCorre-style piecewise translation fields, projection-based rotation estimates, dense SimpleITK-based six-degree-of-freedom volume registration, and sparse point-based rigid-volume registration within one canonical microscopy stack model. The platform uses OMIO to normalize heterogeneous microscope files and to preserve metadata, while optional disk-backed Zarr arrays support chunked, memory-efficient processing of image stacks that exceed available memory or reside on remote storage. ZenReg writes registered images together with shift tables, correlation metrics, summary plots, and machine-readable settings. In synthetic benchmarks with known ground truth, ZenReg recovered global 2D and 3D translations with subpixel accuracy across moderate noise and drift regimes, while high-noise and large-drift tests separated the backend behavior: FFT-based methods failed abruptly once image information or shared support became insufficient, StackReg degraded more gradually under severe noise, and piecewise NoRMCorre improved spatially varying local-motion correction where a single global transform was inadequate. Parallel execution reduced runtime for large time series, and ZenReg provided practical full-volume rigid correction for dense and puncta-rich 3D+t stacks. By coupling a modular, extensible backend architecture to transparent sidecar outputs, ZenReg makes motion correction easier to extend, inspect, share, reproduce, and reuse as part of scientific image analysis.

2
EASI-PASS: An accessible pipeline for linking functional imaging and mRNA profiling

Singh Alvarado, J.; Massengill, C. I.; Stern, J.; Amsalem, O.; Ventura, B. F.; Jang, A.; Cook, S.; Veliche, A.; Sunkavalli, P.; Patel, D.; Colaccino, J.; Evans, K. E.; Wang, Y.; Andermann, M. L.

2026-08-26 neuroscience 10.64898/2026.08.21.746328 medRxiv
Top 0.1%
50.5%
Show abstract

We developed EASI-PASS, a reliable, high-throughput method for estimating the molecular identity of functionally characterized cells by merging live imaging with subsequent fixed-tissue imaging using conventional microscopes. Our method matches the shapes and locations of thousands of densely imaged cells between large (>1 mm2) functional images and a thick, expanded, and cleared EASI-FISH tissue volume to assess gene expression. This approach is more efficient than alignment to thin sections and recovers the molecular identity of ~78% of cells. In acute brain slice imaging from the mouse parabrachial nucleus during optogenetic stimulation of long-range spinal inputs, we observed fine-scale specificity in the molecular identity of spinorecipient neurons. In the awake mouse visual cortex, we observed distinct arousal modulation and spatial falloff in correlations within and across interneuron classes. Thus, EASI-PASS provides reliable and efficient alignment of cellular activity with molecular identity.

3
Fast calcium-dependent fluorescent labeling for recording of neuronal activation

Porzberg, N.; Heck, J.; Wilhelm, J.; Benjaminsen, J.; Bluemel, T.; Huppertz, M.-C.; Noh, K.-M.; Thumberger, T.; Heine, M.; Wittbrodt, J.; Saka, S. K.; Hiblot, J.; Johnsson, K.

2026-08-11 neuroscience 10.64898/2026.08.05.742984 medRxiv
Top 0.1%
46.4%
Show abstract

Calcium transients encode cellular and neuronal activity across timescales ranging from milliseconds to hours, yet linking these transient signals to downstream molecular states remains a major challenge. We recently introduced Caprola, a calcium-dependent protein labeling tool that converts calcium transients into permanent fluorescent marks for later analysis. In this way, Caprola enables tracking of neuronal activities in animal models as well as retrospective identification of labeled cells for isolation and transcriptomic analysis. However, the relatively slow labeling kinetics of Caprola required high concentrations of fluorophore probe and relatively long labeling times, which limits its sensitivity and applicability, in particular in vivo. To address this limitation, we generated Caprola variants with up to 29-fold faster labeling rates than their predecessor. We demonstrate that our new Caprola variants record calcium transients in cells and in zebrafish larval brains under conditions where previous Caprola variants did not show labeling. We further expand the applicability of Caprola to activity-dependent marking of postsynaptic compartments, opening new avenues for coupling functional activity histories with downstream molecular and transcriptomic analyses.

4
AnchorR: A QuPath and R interface for collaborative exploration of spatial transcriptomics and histology

Morris, C. A.; Bastian, W. C.; Cui, Y.; Kurago, Z.; Douglass, E. F.

2026-08-11 bioinformatics 10.64898/2026.08.05.742985 medRxiv
Top 0.1%
45.0%
Show abstract

Single-cell spatial transcriptomics can connect molecular cell states with tissue morphology, but this promise depends on accurate registration to histopathology. In serial sections, however, tissue borders often differ because of sectioning artifacts, staining variability, and field-of-view acquisition, limiting conventional area-based registration. We developed AnchorR, an expert-guided workflow for coarse-grained alignment of hematoxylin and eosin (H&E) images with CosMx Spatial Molecular Imaging data. Bioinformaticians first define and color-code cell types in Seurat, and pathologists then identify corresponding internal landmarks using QuPath overlays. AnchorR combines these paired landmarks to estimate affine transformations, quantify residual error, and support visual quality control and anchor refinement. Using six oral pre-cancerous tissue sections, we identified 60 cross-modal landmarks. Fitting each section independently reduced mean landmark error from 121.5 {micro}m with a single whole-slide transformation to 14.6 {micro}m. Cross-validation further showed that increasing the number of anchors improved robustness, with nine-anchor fits achieving approximately 20 {micro}m error, or about one cell diameter. AnchorR is designed to complement automated computer-vision methods by providing reliable tissue-level alignment when border mismatch makes global registration difficult. By creating a shared workspace for pathologists and bioinformaticians, it operationalizes an expert-in-the-loop approach and makes feature-based multimodal registration accessible without specialized computer-vision expertise or high-performance computing.

5
FOCUS-3D: Robust, generalizable volumetric cell segmentation for three-dimensional fluorescence microscopy

Zhang, Q.; Mu, Z.; Liu, B.; Chi, Y.; Li, D.; Wang, W.; Ni, J.-Q.; Wan, Y.; Yu, L.; Navajas Acedo, J.; Yu, G.

2026-08-28 bioinformatics 10.64898/2026.08.25.746907 medRxiv
Top 0.1%
39.7%
Show abstract

Understanding how cells establish spatial organization within tissues is a fundamental question in life sciences. While modern three-dimensional fluorescence microscopy captures large-volume tissue architecture, extracting quantitative cellular insights from complex volumetric datasets remains a major barrier. Here, we introduce FOCUS-3D, a robust, broadly generalizable volumetric cell segmentation framework built on a large, diverse manually annotated cell resource and advanced AI designs. Integrating volumetric representation learning, multi-scale feature extraction, and query-based mask prediction, FOCUS-3D achieves state-of-the-art performance across diverse species, tissues, fluorescent reporters and imaging modalities. During zebrafish (Danio rerio) development, FOCUS-3D uncovers three successive phases of notochord morphogenesis. We disentangle early motility-driven rearrangements from later cell shape remodeling and tissue repacking, and further link these morphological states to spatial and developmental transcriptional programs across independent datasets.

6
Structural-functional calibration corrects single-neuron identity errors in volumetric calcium imaging

Liu, X.; Gou, D.; Song, C.; Zhao, J.; Liu, M.; Rao, S.; Liang, Y.; Xu, L.; Mao, H.; Liu, Y.; Wang, J.; Ma, L.; Li, H.; Guo, C.; Chen, L.

2026-08-24 neuroscience 10.64898/2026.08.20.745679 medRxiv
Top 0.1%
34.1%
Show abstract

Volumetric calcium imaging is increasingly used to capture larger neuronal populations at higher throughput, but high-speed axial sampling can compromise single-neuron identity. Here we identify cross-plane identity duplication as a structured error in volumetric imaging: anisotropic axial blurring and plane-wise functional segmentation can repeatedly detect the same neuron across adjacent planes, creating duplicate functional nodes that inflate neuronal counts and distort network phenotypes. We developed Comprehensive Label-Guided (CLG) volumetric imaging, a structural-functional calibration framework that uses nuclear labels as stable three-dimensional identity anchors for calcium signals. CLG combines nuclear labeling, deep-learning-based 3D segmentation, anatomical registration and identity-guided trace reassignment. In larval zebrafish whole-brain recordings, CLG resolved ~30,000 redundant detections and reduced estimated neuronal counts by 37-46%. In mouse visual cortex, CLG consolidated ~40% of putative duplicates and recovered over 2,000 active neurons missed by calcium-only analysis. Across baseline and perturbed conditions, calibration stabilized graph-derived measurements of hub organization, long-range correlations and network resilience. CLG therefore defines an anatomy-constrained identity-calibration layer for reliable single-neuron-resolved volumetric imaging.

7
scATrans: annotating single-cell differential expression as transcription- or stabilization-weighted using unspliced RNA

Li, Z.; James, A.; Li, S.

2026-08-07 bioinformatics 10.64898/2026.08.03.740741 medRxiv
Top 0.1%
33.6%
Show abstract

Single-cell differential expression (DE) reports changes in mature mRNA abundance, but the same fold-change can reflect faster synthesis or slower decay. Metabolic labeling resolves this ambiguity but is costly and cannot be applied retrospectively to the vast majority of published scRNA-seq. scATrans closes this gap using layers every standard pipeline already generates: from DE-selected genes, a reference-corrected unspliced residual annotates each change as transcription- or stabilization-weighted, with no additional experiment. Benchmarked against metabolic-labeling systems with independent kinetic ground truth, the residual separates the two mechanisms at matched mature abundance (ROC-AUC 0.68-0.74, full-length NASC-seq2 K562; 0.59-0.63, 3' scEU-seq RPE1); effect size scales with intron capture, not model complexity, and explicit kinetic fitting adds nothing over the static contrast. Per-gene, the residual recovers the classical bulk exon-intron contrast (EISA); what scATrans adds is the inference layer single-cell reanalysis actually needs--DE-defined membership, gene-structure correction, a capture-regime reliability pre-flight, induction-matched testing, and a permutation-calibrated program score--so that confident calls are reserved for where the data support them: gene programs, not single genes. Applied to standard 10x data with no labeling, scATrans recovers textbook post-transcriptional biology: a curated AU-rich-element program is called stabilization-weighted in LPS-stimulated PBMCs (confirmed by per-donor pseudobulk DE in an independent four-donor cohort), while a glucocorticoid-response program is called transcription-weighted in dexamethasone-treated A549 cells-- opposite mechanisms recovered from unlabeled counts. scATrans turns any spliced/unspliced-resolved DE table into a mechanism-typed one, retrospectively and at scale. Availability and implementationscATrans requires Python [≥]3.9, interoperates with the scverse ecosystem (AnnData, Scanpy), and is released under the Apache-2.0 license. Install with pip install scatrans or from Bioconda. Documentation and tutorials: https://scatrans.readthedocs.io. Analyses in this manuscript use software version 0.10.9.

8
Vipsania: Unsupervised Deep Gene Finding

Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.

2026-08-30 bioinformatics 10.64898/2026.08.26.747235 medRxiv
Top 0.1%
32.4%
Show abstract

Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.

9
Meso2EM: a cross-scale CLEM workflow linking mesoscale functional imaging to targeted electron microscopy

Oomoto, I.; Murate, M.; Sohn, J.; Tamura, M.; Hatada, S.; Egawa, N.; Odagawa, M.; Suga, M.; Kawaguchi, Y.; Murayama, M.; Kubota, Y.

2026-09-01 neuroscience 10.64898/2026.08.25.746890 medRxiv
Top 0.2%
30.6%
Show abstract

Meso2EM is a correlative light and electron microscopy workflow that transfers neurons selected from mesoscale functional images to targeted electron microscopy. We recorded Ca{superscript 2} signals from layer 2/3 neurons across a contiguous 3 x 3 mm cortical field in awake mice and reidentified a selected neuron after fixation and tangential sectioning. Lectin-labeled vascular architecture served as a shared landmark across in vivo two-photon imaging, confocal microscopy, laboratory micro-CT of resin-embedded tissue, and block-surface scanning electron microscopy, guiding focused-ion-beam scanning electron microscopy to the target cell body. The same progressive-targeting principle also supported serial ATUM-SEM reconstruction of an in vivo-tracked dendrite and serial transmission electron microscopy of optically selected dendrites from a patch-clamp-recorded Martinotti cell. Meso2EM therefore provides a practical route for preserving target identity across large changes in scale and specimen state while restricting electron-microscopy acquisition to a selected region.

10
ACCREDIT: A Quality-Aware Agentic Engine for Cell-resolved Cross-modal Image Registration with Dynamic Iterative Tuning

Zhou, L.; Zhao, F.; Ren, T.; Goodyear, S. M.; Tang, C.; Li, B.; Zhang, T.; Chen, Y.; Sears, R. C.; Mills, G. B.; Kardosh, A.; Xia, Z.

2026-08-14 bioinformatics 10.64898/2026.08.08.743602 medRxiv
Top 0.2%
30.1%
Show abstract

Spatial omics across complementary modalities is transforming our understanding of tissue architecture. Realizing this potential requires accurate and robust registration of cross-platform molecular images with hematoxylin-and-eosin (H&E) sections, the primary morphological reference for pathology. Existing methods, however, often fail silently when image orientation is unknown, image contrast is inverted, or tissue overlap is incomplete, producing erroneous registrations without alerting users or attempting recovery. Here, we present ACCREDIT, a quality-aware agentic framework that redefines cross-modal registration as an adaptive decision-making process rather than a one-shot computation. ACCREDIT combines deterministic registration pipelines with a reference-free composite quality score that automatically evaluates registration quality and rejects plausible but biologically incorrect registrations. When registration quality is insufficient, a large language model (LLM)-based rescue agent autonomously diagnoses failure modes and selects targeted recovery strategies, while an optional strategy-learning module captures expert-validated corrections for future reuse. Across Xenium, CODEX, cell-boundary, and IHC-to-H&E registration tasks, ACCREDIT outperformed competing methods by detecting registration failures and improving alignment quality through automated recovery and rescue. Ultimately, ACCREDIT enables robust integration of histology and spatial molecular profiling, providing a foundation for translating spatial omics into routine H&E-based pathology workflows.

11
ZEISS arivis Cloud: a cloud-based platform for deep learning model training and scalable bioimage analysis

Bhattiprolu, S.; Toor, M.; Soyer, S.

2026-08-13 bioinformatics 10.64898/2026.08.07.743540 medRxiv
Top 0.2%
28.8%
Show abstract

Modern biological imaging generates large, complex datasets that require scalable and reproducible image analysis methods. Deep learning has demonstrated strong performance on bioimage segmentation tasks, but training custom models has remained inaccessible to many researchers due to requirements for GPU infrastructure, programming expertise, and large annotated training datasets. ZEISS arivis Cloud is a browser-based platform for deep learning model training that addresses these barriers through partial annotation support, AI-assisted labeling with SAM (Segment Anything Model), pretrained model initialization, and automatically configured training pipelines requiring no machine learning expertise. The platform supports two segmentation tasks: semantic segmentation using a U-Net-style architecture with an EfficientNet encoder and PixelShuffle decoder, and instance segmentation based on Mask2Former with a Swin-Tiny backbone. Both pipelines incorporate microscopy-specific adaptations including smooth tiling, multi-channel input support, dataset-specific normalization, and partial-annotation-aware loss functions protected by patents US-20240078681-A1 and US-20250111519-A1. Trained models integrate directly with ZEISS arivis Pro for pipeline-based image analysis, ZEISS arivis Hub for parallel execution across large datasets, and ZEISS ZEN for content-aware guided acquisition. We describe the platform architecture, training methodology, segmentation architectures, reproducibility and versioning mechanisms, and FAIR compliance, and illustrate the complete workflow through two intestinal organoid imaging examples. arivis Cloud is freely accessible to student users; other users access the platform via subscription at https://www.arivis.cloud/.

12
GIAnT: a Glutamate Imaging Analysis Toolbox

Xie, M. E.; Friedrich, J.; Wirsching, E.; Shibu, C. J.; Seyedolmohadesin, M.; Ouellette, N.; Wang, T.; Svoboda, K.; Charles, A. S.; Podgorski, K.

2026-08-13 neuroscience 10.64898/2026.08.07.743580 medRxiv
Top 0.2%
28.4%
Show abstract

Recent advances in fluorescent indicators and optical microscopy now enable in vivo synaptic imaging of glutamate, which transmits the majority of signals between neurons in the brain. Extracting fluorescence signals from these recordings is complicated by the minuscule scale and dense clustering of synapses on dendrites, as well as brain motion in behaving animals. Here we present the Glutamate Imaging Analysis Toolbox (GIAnT), a set of automated tools for glutamate imaging data that corrects sample motion, identifies active synapses with super-resolution precision, and extracts synaptic fluorescence signals. Compared to methods designed for cellular imaging, GIAnT reduces motion artifacts, more accurately identifies active synapses, and improves extracted signal quality by reducing contamination from overlapping synapses. By pairing in vivo glutamate imaging with post hoc expansion microscopy, we find that >70% of the putative synapses extracted using GIAnT matched one-to-one with glutamatergic synapses onto the postsynaptic cell. Our results establish GIAnT as an automated and validated pipeline for processing synaptic glutamate imaging data at scale.

13
A multimodal, correlative magnetic tweezers-TIRF platform for high-throughput single-molecule interrogations

Feiz, M. S.; Cnossen, J.; Wubulikasimu, Y.; Quack, S.; Bugea, T.; Zupnik, A.; Prajapati, R. K.; Rakib, A.; Papini, F. S.; Smitskamp, Q.; Malinen, A. M.; Dulin, D.

2026-08-20 biophysics 10.64898/2026.08.11.744120 medRxiv
Top 0.2%
26.5%
Show abstract

Single-molecule techniques can resolve biological reactions at unmatched detail, but their low throughput and single-modality readouts have kept them out of data-intensive pipelines such as omics and drug discovery, and beyond reach of low-yield biological systems. Here we introduce a multimodal platform integrating high-throughput magnetic tweezers with ultra-wide-and flat-field objective-based total internal reflection fluorescence, enabling simultaneous force, torque, multicolor fluorescence, and temperature-dependent measurements on up to thousands of individual molecules in parallel and in real time. We demonstrate accurate single-molecule Forster resonance energy transfer (smFRET) for prism-based spectral imaging, capture temperature-dependent hairpin folding dynamics at high temporal resolution with smFRET and use correlative torque-fluorescence measurements to unravel the open-complex formation dynamics during bacterial transcription initiation. By unifying high resolution, throughput, and multimodal readout, this platform enables multidimensional dissection of complex biomolecular reactions with high statistical confidence, unlocking single-molecule biophysics for integration with drug discovery, omics, and cryo-EM workflows.

14
Discovery and Targeting of a Cryptic Human Proteome

Chick, J. M.; Woodfin, A. R.; Weir, J.; Blanchette, M.; Schwartz, A. S.; Garnar-Wortzel, L.; Polera, C. A.; Steiniger, S. C. J.; Kelly, M.; Wilson, K.; Bass, J. A.; Jaeger, A. M.; Ajjawi, I.; Dambacher, C. M.

2026-08-11 genomics 10.64898/2026.08.05.743124 medRxiv
Top 0.2%
26.5%
Show abstract

First-in-class therapeutics require first-in-class biology. Yet despite decades of genomic and proteomic cataloging, vast regions of the human transcriptome remain dark and their encoded proteins invisible. Here we present RyboCypher, an integrated RNA-sequencing and AI-assisted proteogenomics platform that systematically maps the RyboCypher-derived "dark" transcriptome to unannotated peptides, predicting and empirically identifying cryptic proteins across the uncharted genome. Applied to cancer cell lines, patient tumors, and matched healthy tissues, RyboCypher resolved [~]8.3 million dark RNA isoforms and [~]16 million candidate ORFs. Interrogating these against [~]0.5 billion MS/MS spectra from cellular proteomics, membrane proteomics, and immunopeptidomics datasets (comprising a total of >8,000 raw MS data files ([~]7TB of MS data), derived from 2,229 patient samples), we empirically identified [~]80,000 cryptic peptides ([~]10,000 cancer-associated or cancer-upregulated) at <1% FDR. Altogether, these datasets establish the CypherAtlas, a comprehensive proteogenomic atlas of an unreported proteome comprising thousands of novel proteins, including membrane proteins with targetable extracellular domains, and intracellular proteins accessible through antigen presentation. By linking dark-RNA transcripts, predicted proteins, and patient-level metadata across RyboDyns proprietary experimental data, CypherAtlas further provides the training substrate for multi-modal models such as DarkCypher, which is being developed to prioritize cryptic targets and to forecast their expression in new patient samples. As proof of therapeutic potential, we disclose evidence for a cancer-associated, cryptic protein expressed from the YBX1 locus, (cryptic YBX1; cYBX1) and demonstrate selective in vitro tumor cell killing through a cryptic peptide-MHC (pMHC) complex derived from this protein with a TCR-mimic (TCRm) antibody when formatted as antibody drug conjugates (ADCs). Together, RyboCypher and CypherAtlas establish the dark proteome as a vast and previously inaccessible reservoir of novel targetable biology, laying the foundation for the next generation of first-in-class therapeutics. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=108 SRC="FIGDIR/small/743124v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@1516d25org.highwire.dtl.DTLVardef@d4b796org.highwire.dtl.DTLVardef@160efd8org.highwire.dtl.DTLVardef@1226f70_HPS_FORMAT_FIGEXP M_FIG C_FIG

15
CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop

Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.

2026-08-20 bioinformatics 10.64898/2026.08.15.745007 medRxiv
Top 0.2%
26.4%
Show abstract

Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.

16
Reference-free protein sequencing by consensus assembly of redundant de novo peptide reads

Nilsson, A.; Sporre, E.; Schulte, D.; Snijder, J.; Edfors, F.; Käll, L.

2026-08-19 bioinformatics 10.64898/2026.08.13.744110 medRxiv
Top 0.2%
26.4%
Show abstract

Reading a proteins sequence from tandem mass spectra without a reference is limited by single-spectrum accuracy, most acutely across the hypervariable complementaritydetermining regions of antibodies. Broadly specific proteases tile a protein with long, overlapping peptides, so every residue is covered by many independent de novo reads. borgonovo assembles their per-step probability profiles into a reference-free per-residue consensus, seeding templates from mass-closure-consistent reads and recruiting the rest by substitution- tolerant alignment and per-column voting. Re-decoding each spectrum with a prior from its consensus position lifts amino acid accuracy on placed spectra from 0.80 to 0.87. On the therapeutic antibody trastuzumab, nine proteases cover its heavy and light chains completely at 0.88 fixed-window identity, and 0.93 on the pruned assembly once local indels are accommodated. Applied unchanged to five secretome proteins and trastuzumab with three proteases, it reaches 0.87 mean fixed-window identity over 82% coverage. borgonovo is open source and works with most de novo sequencers, so redundant digestion turns any of them into a protein sequencer where no reference exists.

17
Automating scientific annotations for open transcriptomic profiles via multi-stage agents

Zhang, X.; Paithankar, S.; Pu, J.; Murtaza, M. S.; Shankar, R.; Leshchiner, D.; Koirala, S.; Palmer, Z.; Nault, R.; Li, X.; Xie, Y.; Chen, B.

2026-08-20 bioinformatics 10.64898/2026.08.19.745739 medRxiv
Top 0.2%
25.7%
Show abstract

Public transcriptomic repositories contain millions of samples, yet their large-scale reuse is hindered by heterogeneous and inconsistently reported metadata. In the Gene Expression Omnibus (GEO), key biological information is often distributed across study- and sample-level records, requiring context-dependent interpretation. Here we present GEOMeta, a large language model (LLM)-based multi-stage workflow with task-specialized agents for automated GEO metadata curation. The pipeline separates metadata retrieval, task-specific information extraction, field standardization, ontology mapping and quality control. Using GEOMeta, we generated standardized annotations for approximately 600,000 human bulk RNA-seq samples. To demonstrate its utility, we benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings. We further prospectively annotated newly submitted GEO studies and evaluated 22 frontier LLMs. Recent open-source Flash models achieved annotation quality comparable to leading reasoning models while reducing costs by an order of magnitude. GEOMeta provides a scalable resource and reproducible framework for metadata curation.

18
HI-JEPA: A World Model of Molecular Organization Learned from Measured Proximity

Shihabi, R.; Karmali, S.; Vaughan, B.; Taraman, S.; Kellis, M.

2026-08-20 systems biology 10.64898/2026.08.14.744901 medRxiv
Top 0.2%
25.5%
Show abstract

Proteins act through the company they keep. Which molecules occupy the same nanoscale neighborhood in intact tissue determines what can physically interact, and disease rearranges those neighborhoods before it changes anything a sequence records. That quantity (measured proximity between molecular species in unperturbed tissue) has never been acquired broadly enough to train on. Published colocalization arrives study by study and never accumulates into a graph. The measurement has to be made rather than collected. We built ASCEND, a spatial computing platform that measures pairwise molecular proximity from expansion microscopy at molecular resolution in intact tissue, and applied it to 164 proteins across 37 imaged regions in five studies, spanning cultured neurons, isolated synapses and mouse cortex in disease and control. HI-JEPA is a representation trained on those measurements. Each protein is one embedding, trained to predict the embeddings of its measured neighbors in latent space; it never reconstructs its input and generates no negatives. A set of proteins measured in one neighborhood forms a configuration, which is the object the model perturbs and plans over. The representation performs operations a sequence model cannot. It names a protein from the bare geometry of a microscopy point cloud, matched against 234,048 deposited structures, at top-1 accuracy 0.748 against a chance rate of 1.0 x 10-5. It predicts physical interaction between sequence-dissimilar proteins that were both withheld from training at AUC 0.908, where ESM-C 6B reaches 0.514 against partner-count-matched negatives. It recovers a held-out complex member in the top 100 of 13,447 candidates at recall 0.954, against 0.514 for a ranking built from complex frequency alone. Asked which partners a knockout disrupts, it recovers the experimentally observed ones at recall@100 0.640; asked the same question about a different protein, with the ranking rule and denominators unchanged, it recovers 0.028, so the answer follows the action. Given 5xFAD mouse cortex with no disease label, no reward and no indication that amyloid is relevant, ranking 1,574 measured assemblies by their departure from wild type returns amyloid-{beta} bound to AMPA receptor subunits in nine of the top ten. Planning over the same configurations independently selects the same subunits (GluA2, GluA3, GluA4) and predicts that disrupting the PSD-95 scaffold worsens the configuration, both agreeing in sign with experiments the model never saw. Ablating the measured-proximity channel at training time degrades cross-scale partner recovery from median rank 14 to 68 while leaving navigation and within-scale dynamics intact; ablating the perturbation channel does the reverse. The cross-scale capability therefore comes from the measurement and not from having seen more data. The intended application is target nomination in diseases where sequence and structure supply no starting point. Note: This is a capability report. The architecture, the training procedure and the acquisition protocol are proprietary and are not described. Section 4.2 gives the evaluation protocol behind every number reported.

19
PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics enables robust generative modeling of gene expression and scales single-cell integration to 100 million cells

Wang, N.; Cardenas, C.; Nieto Caballero, V. E.; Turner, D.; Feinberg, H.; Yuan, D.; Scott, N.; DeBerardine, M.; Dan, S.; Caceres, L.; Schembri, J.; Yao, Z.; Lee, C.; Pillow, J. W.; Krienen, F. M.

2026-08-12 bioinformatics 10.64898/2026.08.06.743394 medRxiv
Top 0.3%
22.8%
Show abstract

Single-cell RNA technologies enable the routine acquisition of transcriptomic atlases. However, these molecular profiles are influenced by overlapping sources of variation. Since these covariates confound comparisons, data integration is the first step in most analyses. Three challenges remain: correcting strong batch effects, scaling to millions of cells, and modeling how covariates influence gene expression. To address these challenges, we developed PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics, a deep learning framework whose central feature is a generative model of gene expression data. Additionally, PIANO achieves robust integrations and trains 10x faster than previous methods. PIANO accurately integrates single-cell data across species and across single-cell and spatial transcriptomics modalities. As practical applications, PIANO models spatially-resolved gene expression during Alzheimers disease progression in human brains and integrates over 100 million cancer cells to model drug perturbations. In summary, PIANOs integration and generative modeling capabilities will empower novel insights for countless future studies.

20
cFAR and Relative Signal: Diagnosing Preferred Orientation in Single-Particle Cryo-EM

Peretroukhin, V.; McLean, M.; Punjani, A.

2026-08-18 biophysics 10.64898/2026.08.11.744264 medRxiv
Top 0.3%
22.4%
Show abstract

The quality of single particle cryo-EM reconstructions can be severely degraded when an insufficient variety of 3D particle orientations is present in the image data, limiting downstream model building and interpretation. However, it is often difficult to ascertain whether or not a particular dataset suffers from such preferred orientation since the required orientation coverage depends on target geometry, alignment accuracy, and particle quality. To simplify diagnosis of preferred orientation, we present two complementary methods. First, the conical Fourier Shell Correlation Area Ratio (cFAR) compares the worst- and best-correlating conical regions of 3D Fourier space to quantify half-map anisotropy into a single, easily interpretable score ranging from zero to one. Second, Relative Signal, a companion to cFAR, directly relates signal content to viewing direction so that under-sampled views can be identified. We characterize our methods and compare them to existing anisotropy detection approaches on synthetic data and on 14 real datasets that span sundry molecular weights and structure types. Implementations of both cFAR and Relative Signal are included in CryoSPARC v4.5 and later versions.